Province of Tarlac
Calibrating Long-form Generations from Large Language Models
Huang, Yukun, Liu, Yixin, Thirukovalluru, Raghuveer, Cohan, Arman, Dhingra, Bhuwan
To enhance Large Language Models' (LLMs) reliability, calibration is essential -- the model's assessed confidence scores should align with the actual likelihood of its responses being correct. However, current confidence elicitation methods and calibration metrics typically rely on a binary true/false assessment of response correctness. This approach does not apply to long-form generation, where an answer can be partially correct. Addressing this gap, we introduce a unified calibration framework, in which both the correctness of the LLMs' responses and their associated confidence levels are treated as distributions across a range of scores. Within this framework, we develop three metrics to precisely evaluate LLM calibration and further propose two confidence elicitation methods based on self-consistency and self-evaluation. Our experiments, which include long-form QA and summarization tasks, demonstrate that larger models don't necessarily guarantee better calibration, that calibration performance is found to be metric-dependent, and that self-consistency methods excel in factoid datasets. We also find that calibration can be enhanced through techniques such as fine-tuning, integrating relevant source documents, scaling the temperature, and combining self-consistency with self-evaluation. Lastly, we showcase a practical application of our system: selecting and cascading open-source models and ChatGPT to optimize correctness given a limited API budget. This research not only challenges existing notions of LLM calibration but also offers practical methodologies for improving trustworthiness in long-form generation.
- Asia > Singapore (0.04)
- Asia > Philippines > Luzon > Central Luzon > Province of Tarlac > City of Tarlac (0.04)
- North America > Canada > Ontario > Toronto (0.04)
- (7 more...)
- Health & Medicine > Pharmaceuticals & Biotechnology (0.68)
- Media > Music (0.67)
- Leisure & Entertainment (0.67)
Data Management Assistant ( Stay-in Set-up) at Pilmico Foods Corporation - Tarlac, Philippines
Pilmico Foods Corporation is the integrated agribusiness and food company of Aboitiz Equity Ventures Inc. (AEV). Composed of four divisions: Flour, Feeds & Animal Health, Farms, and Trading, we are well positioned at the beginning of the value chain. True to our brand promise of being Partners for Growth, we nurture our business and communities by providing business solutions and building partnerships for growth. We operate in the Philippines nationwide and have a growing international presence in the following ASEAN countries: Vietnam, Thailand, Indonesia, Malaysia, Myanmar and Hong Kong. Investing in talent and upholding Aboitiz values of Integrity, Teamwork, Innovation, and Responsibility are key drivers to sustaining the growth of our business.
- Asia > Philippines > Luzon > Central Luzon > Province of Tarlac (0.40)
- Asia > Vietnam (0.26)
- Asia > Thailand (0.26)
- (4 more...)